skip to content

cdxgen, Syft and Trivy can all emit CycloneDX - what does a generator's native model change?

level: middleimportance: should knowfreq 45%

answer

  1. the flag picks a serialisation
  2. each tool has its own model
  3. empty fields, not absent components
  4. converters cannot invent data
  5. dependency edges come from the input

basics

~20 s

Every generator keeps its own component model and projects it onto whichever format you ask for. cdxgen models to CycloneDX directly; Syft and Trivy export to either format, so spec fields their model never captured come out empty.

solid answer

~50 s

The format flag chooses a serialisation, not a data source. cdxgen comes from the CycloneDX project and models components against that spec first, so its CycloneDX output is the shape it natively thinks in. Syft has its own internal model and writes SPDX or CycloneDX as projections of it; Trivy likewise emits either format from its own scan-result model. The consequence is that the interesting fields are decided by the generator, not by the format: supplier and author fields, license expressions and especially the dependency graph come out populated only if the tool's model and its input carried them. A filesystem scan of an image usually cannot reconstruct which package required which, so the dependency section is thin in either format. Judge a generator by inspecting the fields your consumer needs in a real document, not by the format name on the flag.

go deeper

for a junior

Know that both SPDX and CycloneDX can carry a component list, that generators offer a flag to choose, and that the choice does not change which components were found.

for a middle

Explain the projection: internal model out to a spec, empty fields where the model had nothing, property bags for tool-specific extras, and why a dependency graph depends on the input rather than the format.

for a senior

Show that you evaluate a generator by diffing real documents field by field against what a downstream consumer reads, instead of comparing feature lists.

for a principal

Own the estate-level call: one format across the fleet is worth a conversion step, but only after you have confirmed the fields your policy and audit consumers depend on survive it.

## Format is a serialisation, not a source of truth The two mainstream SBOM formats have different heritage. **SPDX** grew out of the Linux Foundation with a licence-compliance heritage, so its vocabulary is strong on files, licences, suppliers and relationships. **CycloneDX** grew out of OWASP with a security-analysis heritage, so its vocabulary is strong on components, package URLs, services and attaching security context. Both can carry a component inventory; they emphasise different things and name them differently. A generator sits *behind* that choice. It collects evidence, builds an internal model of components and their relationships, and then serialises. Three practical cases: - **cdxgen** is the CycloneDX project's own generator. It models against that spec directly, which is what "native" means here: there is no translation step between the tool's idea of a component and the document's. - **Syft** keeps its own richer internal model and can emit SPDX, CycloneDX, or that internal model. Neither standard format is privileged; both are projections. - **Trivy** produces its result set from its analyzers and can serialise it into either format too. ## What projection costs you When a model is projected onto a spec, three things happen: **Fields the spec has and the model lacks are empty or synthesised.** If the tool never learned a component's supplier, the supplier field is absent or filled with a placeholder. A consumer checking for supplier data across your estate sees the gap, not the reason for it. **Fields the model has and the spec lacks are pushed into escape hatches.** CycloneDX has a `properties` bag; SPDX has annotations and external references. Tool-specific detail — which cataloger found the component, which path evidenced it — ends up there, which is useful for debugging and invisible to a strict consumer. **Relationships are the hardest thing to project.** A dependency graph requires knowing who required whom. A lockfile carries that; a build-time hook inside a package manager carries it; a scan of an unpacked image filesystem generally does not, because the installed files do not record who asked for them. So a CycloneDX `dependencies` section or an SPDX `DEPENDS_ON` relationship set can be thin or flat regardless of which format you selected. That is a property of the vantage point the generator used, projected through the format. ## Converting between formats Converters exist and produce schema-valid output. What they cannot do is invent data. Converting SPDX to CycloneDX remaps names, splits and merges fields, and drops anything with no target expression. Round-tripping loses a little more each way. So "we generate SPDX and convert for the customer who wants CycloneDX" is workable only if the fields that customer checks survive the trip — verify that on a real document rather than assuming it. ## How to choose in practice Work backwards from the consumer: 1. **List the fields something downstream actually reads.** A policy check that keys on package URLs needs correct purls per ecosystem. A licence audit needs licence expressions. A contract clause may name a format explicitly. 2. **Generate one document per candidate tool over a representative artifact** and diff the fields, not the totals. 3. **Pick the generator whose vantage point sees your components at all** — a polyglot repository with several build systems needs a generator that recognises each of them, and that constraint usually decides the tool before the format does. 4. **Standardise the format afterwards**, because a fleet of documents in one format is far cheaper to query than a mix, even if some tools then need a conversion step. ## The interview point Candidates often answer "CycloneDX is for security, SPDX is for licences, pick one". That is the heritage, not the decision. The stronger answer is that the document you get is bounded by what the generator saw and modelled, and the format only decides how that is written down.

  • Is converting an SPDX document to CycloneDX equivalent to generating CycloneDX directly?
    Only for the data that survives. A converter remaps and drops fields; it never recovers something the source document did not carry, so a supplier or relationship the original generator never learned stays missing. Generating natively can populate more, because the tool still has its evidence in hand. Test it: convert, then diff the fields your consumer reads against a natively generated document.
  • Two generators run over the same project both claim a complete CycloneDX document, yet one has a dependency graph and the other does not. What explains it?
    Vantage point, not format. A generator that read the lockfile or hooked into the build knows which package required which, so it can emit edges. One that inventoried installed files sees a flat set with no requirement information to record. Both documents are valid CycloneDX; only one can answer whether a component is a direct or transitive dependency.
  • Your consumer needs an exact identifier per component. What do you check in the output?
    Whether every component carries a correct package URL for its ecosystem, with namespace, name and version populated, and qualifiers such as architecture where the ecosystem needs them. Check a sample by hand against the real packages. Synthesised CPE strings are the weaker signal, since generators often guess them by pattern, which produces both misses and false matches downstream.

saying these in an interview costs you the question

  • Treats the format flag as if it changed what was collected
  • Says CycloneDX cannot express licences or SPDX cannot express security data
  • Assumes conversion between formats is lossless
  • Blames the format for a missing dependency graph

context