skip to content

Component Identity Matching

purl coordinates, CPE strings, SWID tags and file hashes name one library four ways, and a fork or a distro rebuild breaks the match. Interviewers ask because it explains most bad scan output.

on this pageshow

questions

4

In an SBOM, how does a purl differ from a CPE as a component identifier?

level: juniorimportance: must knowfreq 70%

answer

  1. one is computed, one is looked up
  2. ecosystem coordinates versus a dictionary entry
  3. pkg:npm/... versus cpe:2.3:a:...
  4. not every package has a dictionary entry
  5. a component can carry both, plus hashes

basics

~20 s

A purl is a coordinate computed from where a package actually came from: ecosystem, namespace, name, version. A CPE is a string from a curated dictionary that a human assigns, so it is often missing, duplicated or ambiguous.

solid answer

~40 s

A package URL (purl) encodes the package manager's own coordinates: `pkg:type/namespace/name@version?qualifiers`, for example `pkg:npm/[email protected]` or `pkg:maven/com.fasterxml.jackson.core/[email protected]`. Any build can compute it deterministically from the resolved dependency, which is why SBOM generators emit it. A CPE is a Common Platform Enumeration string such as `cpe:2.3:a:vendor:product:1.2.3:*:*:*:*:*:*:*`, drawn from a dictionary someone else curates; there is no algorithm that turns a package into its CPE, the vendor and product fields are whatever the dictionary author chose, and plenty of packages have no entry at all. Practically: purl is precise for anything installed by a package manager, CPE is the historical key for operating systems, appliances and commercial products. Good SBOM formats carry both plus hashes on the same component, because different data sources join on different keys.

go deeper

for a junior

Be ready to show the shape of each string and say where it comes from: a purl is computed from the package manager, a CPE is looked up in a dictionary. Naming one component that has neither also scores well.

for a middle

Explain why derivability matters: two builds compute the same purl, while two people can produce two different CPEs for one product. Be able to say which scheme covers operating systems and appliances and which covers language packages.

for a senior

Show what you do with components that carry no usable identifier at all: hashes, supplier-plus-version text, and a named gap in the inventory rather than a silent omission.

for a principal

Own the estate-wide decision about join keys and who maintains the mappings between them, and be honest about the classes of asset that will only ever be named by hand and what that costs.

## Why identity is the whole game An SBOM tells you what is inside an artifact. That inventory is only useful if the names in it are the same names used by whoever publishes information about those components. If your inventory says `openssl 1.1.1n` and the authority you consult says `OpenSSL Project / openssl / 1.1.1n`, a join has to happen, and every mismatch is either a false alarm you waste an engineer on or a silent miss you never see. Component identity is the plumbing under that join. ## purl: identity you can compute A package URL is a compact, URL-shaped coordinate: ``` pkg:type/namespace/name@version?qualifiers#subpath ``` - `type` is the ecosystem: `npm`, `maven`, `pypi`, `golang`, `cargo`, `gem`, `nuget`, `deb`, `rpm`, `oci`. - `namespace` is the ecosystem's grouping concept: an npm scope, a Maven groupId, a distro name for `deb`. - `name` and `version` are that ecosystem's own strings, not a normalised guess. - `qualifiers` carry the rest of what makes the artifact concrete: `arch`, `distro`, `repository_url`, a Maven classifier. The important property is that a purl is **derivable**. The build already resolved that dependency, so the tool that produced the artifact knows the type, namespace, name and version without asking anybody. Two independent builds of the same dependency produce the same purl. That is what makes purl the default identifier in modern SBOM generation. Its limits are equally structural. A purl only exists for something a package manager installed. Vendored source copied into your repository, a statically linked C library, a binary blob dropped into an image, a component someone shaded into another package: none of those have a purl unless a human writes one. ## CPE: identity you look up CPE is the older scheme, designed for enumerating platforms rather than packages. A CPE 2.3 formatted string is thirteen colon-separated fields: ``` cpe:2.3:part:vendor:product:version:update:edition:language:sw_edition:target_sw:target_hw:other ``` `part` is `a` for application, `o` for operating system, `h` for hardware; `*` means ANY and `-` means not-applicable. The `vendor` and `product` fields are dictionary values, and that is the crux: they were chosen by whoever created the entry. The same project can appear under more than one vendor string over its life, a rename produces a second lineage, and a package published this morning may have no entry at all. There is no function from a package to its CPE — matching implementations approximate one with heuristics, and the heuristics are where both false positives and false negatives are born. CPE earns its place where purl cannot reach: operating systems, firmware, appliances and commercial products that are not distributed by any package manager. Those are exactly the things a purl has no type for. ## SWID and hashes: the other two answers A SWID tag (ISO/IEC 19770-2) is an XML document that an installer drops on the machine describing what was installed, with a tag identifier and entities naming the software creator and the tag creator. Its provenance is different from both of the above: purl comes from the build, CPE from a dictionary, a SWID tag from the **installation**. That makes it the natural key for asset management on endpoints and an awkward one for build-time inventory. A cryptographic hash is the one identifier nobody can argue with: it names exactly the bytes you have. Both major SBOM formats carry hashes per component. The catch is that almost nothing publishes vulnerability data keyed by hash, so a hash proves *which* file you hold without telling you *what* it is. It is the tie-breaker, not the join key. ## Carrying more than one SBOM formats are deliberately built to hold several identities for one component. A CycloneDX component has `purl`, `cpe`, `swid` and `hashes` fields side by side. SPDX attaches `externalRefs` to a package, with a PACKAGE-MANAGER category reference for the purl and a SECURITY category reference for the CPE. The NTIA minimum elements ask for supplier name, component name, version and *other unique identifiers* precisely because one scheme is never enough. So the practical answer to "which one" is: emit the purl always, because it is computable and exact; carry the CPE where it exists, because some data sources still key on it; keep hashes so you can prove identity when the names fail; and treat any component that has none of them as a known gap in the inventory, named as such, rather than as an absence of risk.

  • Where does a SWID tag fit alongside purl and CPE?
    A SWID tag is emitted by the installer at install time, so it describes what landed on a machine rather than what a build resolved. That makes it a good key for endpoint asset management and a poor one for build-time inventory: the same component can have a purl from the build and a SWID tag on the device, with nothing but the product name and version linking the two. If you run both worlds, you need an explicit mapping and you should decide which side is authoritative.
  • If a component carries both a purl and a CPE, which should your inventory join on?
    Join on the purl wherever the data source you are consulting speaks purl, because it is exact and reproducible. Keep the CPE for sources that only key on dictionary strings, typically operating systems, firmware and commercial products. Store both rather than picking one globally; a single join key across an estate that contains language packages, OS packages and appliances is the thing that quietly loses components.
  • What identifies a component that has no package manager behind it at all?
    Supplier plus product plus version as text, and a hash of the file you actually hold. That combination is weak for automated matching but it is honest, and it lets a human resolve the component later. Record it explicitly as an identified-by-hand entry so the gap is visible; an inventory that silently omits vendored source or a statically linked library is worse than one that flags it as unresolved.

A purl is a return address printed automatically by the machine that shipped the parcel. A CPE is a name in a directory somebody else maintains, and not everyone is listed under the spelling you expect.

saying these in an interview costs you the question

  • Thinks purl and CPE are two spellings of one name
  • Assumes every package has a CPE entry
  • Believes matching on the component name string is enough
  • Treats a hash as something advisories can be keyed on
  • Says the SBOM format itself determines the identifier

context

open as a page

Why can a distro-rebuilt package version like 1.1.1n-0+deb11u3 produce a false vulnerability match?

level: middleimportance: should knowfreq 52%

basics

~10 s

Distributions backport security fixes into an older upstream version and mark the rebuild with their own revision suffix. A matcher reading only the upstream part flags a package whose fix is already in.

open as a page

Your build shades a JSON parser into your own namespace — how do you keep it visible in your SBOM?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Generate the SBOM from the resolved dependency graph at build time, before relocation erases the coordinates, and ship it with the artifact. Once classes are relocated into your namespace, no consumer can recover the parser's identity by inspecting what you published.

open as a page

A vendor ships an appliance as an RPM under its own product name with no component list — what do you do?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

You cannot compute identity for someone else's build, so the levers are contractual: require a component-level SBOM and a product-keyed advisory channel as purchase terms. Until then, record one opaque component and spend on containment.

open as a page