skip to content

A supplier reuses one CycloneDX serialNumber across three releases. How should your estate key records?

level: middleimportance: should knowfreq 45%

answer

  1. the supplier's identifier is a claim
  2. key on what you can verify
  3. digest of the described artifact
  4. append rows, never upsert
  5. collision is a quality signal

basics

~20 s

Key on the identity you can verify yourself: the digest of the artifact the document describes. Treat supplier-supplied identifiers as metadata, store ingests append-only, and flag a repeated identifier over different subjects instead of overwriting the earlier record.

solid answer

~50 s

A CycloneDX `serialNumber` is meant to identify one BOM, with its `version` field incrementing when that same BOM is revised; an SPDX document namespace is meant to be unique per document and never reused. Both are assertions by whoever wrote the document, and a supplier that reuses a serialNumber across three different product releases breaks the assumption your store depends on. If you upsert on that identifier, the July document silently replaces the March one and you lose the record of what you believed you were running in the spring. The fix is to key each record on the digest of the subject artifact — something you can compute and check — and to carry the document's own identifier as an attribute alongside it. Write ingests append-only, keep the ingestion timestamp, and treat 'same serialNumber, different subject' as a data-quality alert rather than an update.

code

json · 16 lines
json
[
  {
    "bomFormat": "CycloneDX",
    "serialNumber": "urn:uuid:8f1ab2c4-...-9d2e",
    "version": 1,
    "metadata": { "component": { "name": "vendor-telemetry-agent", "version": "4.2.0" } },
    "components": [ "..." ]
  },
  {
    "bomFormat": "CycloneDX",
    "serialNumber": "urn:uuid:8f1ab2c4-...-9d2e",
    "version": 1,
    "metadata": { "component": { "name": "vendor-telemetry-agent", "version": "5.0.0" } },
    "components": [ "..." ]
  }
]

go deeper

for a junior

Recall that a document identifier is written by whoever produced the document, so it can repeat. Know that a store should add records rather than overwrite them.

for a middle

Explain how the identifier is meant to work, what an upsert on it destroys, and why the digest of the described artifact is the safer key. Be able to describe the collision check at ingestion.

for a senior

Demonstrate that you treat a repeated identifier as evidence about the supplier's generation pipeline, and that you can design a record shape which keeps audit history while still answering current-state queries quickly.

for a principal

Own the wider rule for estates assembled from other people's documents: your keys must be computable, their assertions are evidence. Be ready to argue what you do when a strategic supplier's documents are structurally unkeyable.

## Two identifiers, two intentions Inventory formats carry a document-level identifier. In CycloneDX that is the `serialNumber`, a UUID URN whose purpose is to identify **this BOM**; the separate `version` field increments when the same BOM is revised — a correction, an added component, a refreshed licence field — while still describing the same subject. In SPDX the equivalent is the document namespace, a URI that the specification requires to be unique to that document and not reused when a new version of the document is produced. Both are **claims made by the producer**. Nothing enforces them across organisational boundaries, and suppliers routinely get them wrong: a template with a hard-coded UUID, a generator run from the same configuration each release, a manual export copied between products. ## The failure this causes in an estate An estate that treats the document identifier as its primary key is doing an upsert on a value it does not control. When a supplier ships three releases under one serialNumber, three distinct facts about three distinct artifacts collapse into one row. Concretely, you lose: - **Audit truth.** 'What did we believe we were running in March' becomes unanswerable, because the March record was overwritten in July. For a regulated estate, that is often the whole point of keeping the documents. - **Version-aware querying.** The one surviving row describes whichever release happened to land last, so a query about the release you actually deploy silently answers about a different one. - **The signal itself.** A collision is diagnostic — it tells you the supplier's generation pipeline is templated rather than per-build, which colours how much weight the rest of the document deserves. Upserting throws that signal away. ## Key on what you can verify The stable identity in this picture is not in the document header — it is the **subject**: the artifact the document is about. If you can compute a digest over the delivered image, package or firmware blob, that digest is your key, because it is derived from bytes you hold rather than asserted by a party you do not control. The natural record shape is therefore: | Field | Role | |---|---| | subject digest | primary key — the artifact this document describes | | document identifier and revision | attributes, useful for correlation and for detecting supplier problems | | document creation time (stated by producer) | the age of the claim | | ingestion time | when you learned it | | raw document | stored verbatim for audit and reprocessing | When the subject digest is genuinely unavailable — a supplier delivers a document about a product release you cannot hash, such as a device firmware image you never receive as a file — fall back to a **composite** key of supplier identity plus product plus release version, and record explicitly that the key is asserted rather than verified. That distinction should be visible in query results, because it changes how much a downstream decision can lean on the row. ## Append, never upsert Ingestion should add rows, not replace them. Each row is 'at this time, this party told us this about this subject'. A revised document for the same subject is a new row that supersedes the old one for query purposes while remaining readable for audit. Queries then run against the latest non-superseded row per subject, and history stays intact. This also gives you a clean way to handle contradictions. Two documents about the same subject digest that disagree on the component list are not a merge problem; they are a supplier-quality event. Keep both, mark the conflict, and let the query surface it rather than picking a winner silently. ## Detecting the collision The check is small and worth wiring into ingestion: if an incoming document's identifier already exists in the store but the subject digest differs, do not treat it as a revision. Store it, raise a data-quality flag against that supplier, and — where you have the relationship to do so — feed it back so their generator gets fixed. The same check catches the mirror-image mistake inside your own build system, where a templated pipeline stamps every service's document with one identifier. The general principle generalises past this format detail: **in an estate you assemble from documents other people wrote, your keys must be things you can compute, and everything they assert is evidence to be checked rather than schema to be trusted.**

  • What do you key on when you never receive the artifact and cannot compute a digest?
    Fall back to a composite of supplier identity, product and release version, and mark the key as asserted rather than verified. Query results should carry that distinction, because a row keyed on someone else's version string is a weaker basis for a decision than one keyed on bytes you hashed.
  • Two documents describe the same artifact digest but list different components. How do you resolve it?
    Do not merge them. Keep both rows, mark the conflict, and let queries surface it. A disagreement about one subject is a supplier or generator quality event, and picking a winner silently hides exactly the fact you would want to escalate.

saying these in an interview costs you the question

  • Upserts records on the supplier's document identifier
  • Assumes a document identifier is globally unique and honest
  • Overwrites the older record and loses audit history
  • Merges two conflicting documents into one row
  • Keys estate records on a mutable image tag

context