skip to content

What do the OSV, NVD and GHSA advisory databases each assert about a vulnerable package?

level: juniorimportance: must knowfreq 62%

answer

  1. three layers, not three copies
  2. who names it, who enriches it, who aggregates
  3. ecosystem name plus version range
  4. aliases link one flaw's many names
  5. a range claim, not a risk claim

basics

~20 s

NVD enriches a CVE record with a severity score and product applicability. GHSA describes the flaw per package ecosystem with a package name and affected version range. OSV is a schema and aggregator normalising many feeds into ecosystem-native ranges.

solid answer

~50 s

They sit at different points in the same chain. A CVE Numbering Authority assigns the identifier and publishes a record that names the flaw; NVD is the enriched view of that corpus, adding a severity vector and an applicability statement written in vendor-product terms, which fits installed software better than a library published under a package-manager name. GHSA, the GitHub Advisory Database, is ecosystem-native: it names the package the way npm, Maven, PyPI or crates.io names it and gives an affected range and a fixed version in that ecosystem's version scheme, so it matches your lockfile directly. OSV is an open schema plus a service aggregating many feeds into it, expressing affected sets as `introduced`/`fixed` events per ecosystem and carrying an `aliases` list so the same flaw's several identifiers collapse to one. None of the three tells you whether your code calls the vulnerable function.

go deeper

for a junior

Be ready to say in one breath which source names a flaw, which enriches it with a score and applicability, and which aggregates feeds into ecosystem-native ranges. Know that one flaw can have several identifiers.

for a middle

Explain how an affected range is actually represented - introduced and fixed events over an ecosystem's version scheme - and why vendor-product applicability data fits installed software better than a published library.

for a senior

Show that you choose a source per ecosystem and know which feed each tool in your pipeline reads, because that choice, not tool quality, drives most of what your reports say.

for a principal

Own the position that advisory data is third-party input of variable quality: decide which source is authoritative for the estate, budget for the normalisation work, and be able to defend that choice to an auditor.

An advisory database is not one thing. Three distinct jobs are being done in this space, and OSV, NVD and GHSA sit at different points in the chain. Confusing them is why two teams can look at the same service and disagree about what is known. ## Identification CVE Numbering Authorities (CNAs) - vendors, open-source projects and coordinating bodies authorised by the CVE Program - assign an identifier to a reported flaw and publish a record with a description and references. The identifier's job is to give everyone the same name for the same bug. On its own it is not a machine-matchable statement about which versions are affected; the record's prose may say *before 2.14.1* and nothing in the data model forces that to be structured, correct or complete. ## Enrichment NVD is the enriched view of that corpus. It adds a severity vector, a weakness classification and an applicability statement expressed against a dictionary of vendor/product/version entries, which is what lets a tool ask *is this installed product inside the affected set*. Two consequences matter in practice. First, enrichment is a separate step from publication and it can lag: a flaw can be public, with an assigned identifier and a working exploit discussed in the open, while the structured applicability data that scanners key off is still absent. A team whose only source is the enriched national record sees nothing during that window. Second, the applicability vocabulary is written in vendor and product terms. That fits an installed operating system component well; it fits a library published under a package-manager name less well, and the translation between the two is a known source of both misses and over-matches. ## Ecosystem-native curation GHSA is the GitHub Advisory Database. Its entries are scoped to a package ecosystem - npm, Maven, PyPI, NuGet, RubyGems, Go, crates.io and others - and an entry names the package as that package manager names it, with an affected version range and a fixed version in that ecosystem's own version scheme. Entries divide into curated, reviewed advisories where exactly this package-and-range data has been checked, and unreviewed entries imported from the central feed without that curation. The reviewed ones are directly actionable: they can be matched against a lockfile mechanically, with no vendor-to-package translation step in the middle. Most language ecosystems have an equivalent native database of their own - Rust's advisory database, the Go vulnerability database, and per-distribution feeds for OS packages. ## Aggregation and normalisation OSV is two things: an open schema for advisories, and a service that aggregates many source feeds into that schema. The schema is deliberately ecosystem-native. Each entry carries an `affected` list; each affected element names an ecosystem and a package and expresses versions as ranges built from ordered events - an `introduced` version and a `fixed` or `last_affected` version - so a consumer decides membership by evaluating the range rather than by parsing English. Entries also carry an `aliases` list linking the same flaw's identifiers across databases, and a `withdrawn` timestamp for records that have been retracted. ## Aliases: one flaw, several names The same bug routinely has a central identifier, an ecosystem advisory identifier and sometimes a distribution advisory identifier. These are names for one flaw, not three flaws. Counting rows rather than normalised flaws is one of the most common ways a finding count gets inflated, and it is the first thing to check when two reports disagree on volume. ## What none of them assert An advisory asserts that a named package, in a named version range, contains a flaw, and usually which version fixed it. It does not assert that your service calls the vulnerable code, that the package ships in something you run, or that an attacker can reach it. It does not assert that the range is right - ranges are written by people and are sometimes too wide, occasionally too narrow, and can be corrected after publication. Treating an advisory hit as a statement about your risk, rather than as a statement about a version range, is the single most common junior mistake here. ## The practical rule For a language dependency, prefer the ecosystem-native record, because its ranges are expressed in the same versioning language as your lockfile. For an installed operating system package, prefer the distribution's own feed, because only the distribution knows what it patched. Whichever you pick, normalise by alias before you count anything, and know which feed each of your tools is reading - that single fact explains most disagreements.

  • If one flaw carries both a CVE identifier and a GHSA identifier, is that two vulnerabilities?
    No. They are aliases for the same flaw, and the aggregating schema records that relationship explicitly. Normalise by the alias set before counting, otherwise the same bug is reported once per database your tooling reads and the total is meaningless.
  • Why do teams prefer an ecosystem-native advisory over the enriched national record for a language dependency?
    Because the range is expressed in the package manager's own name and version scheme, so matching against a lockfile entry is mechanical rather than a translation from vendor-product terms. Curation of that field is also usually faster than central enrichment, so the actionable data exists earlier.
  • What does an advisory hit still not tell you?
    Whether the vulnerable code is actually called from your service, whether the package is in something you ship or only a build-time helper, and whether the stated range is accurate. It asserts that a version range contains a flaw - everything after that is your judgment.

One earthquake, three records: a catalogue number, a national report adding magnitude and a damage map, and an aggregator that restates every local bulletin in one format.

saying these in an interview costs you the question

  • Says a CVE identifier and a GHSA identifier are always different bugs
  • Thinks the national database assigns the identifiers
  • Treats a missing severity score as meaning no known flaw
  • Assumes every advisory database covers every ecosystem equally
  • Reads an advisory hit as proof the service is exploitable

context