Where does GitHub's dependency graph get its data, and which dependencies does it miss?
answer
- Parsing, not building
- Lock file present or absent changes everything
- Copied-in code declares nothing
- There is an API for what the build knows
basics
~20 sGitHub builds the dependency graph by parsing supported manifest and lock files on the default branch. It misses vendored source, dependencies resolved only during the build, and anything undeclared — gaps you close by pushing a resolved graph through the dependency submission API.
solid answer
~50 sThe graph is **static parsing, not observation**. GitHub reads supported manifest and lock files committed to the default branch — `package-lock.json`, `go.sum`, `Cargo.lock`, `pom.xml`, `requirements.txt` and friends — and builds a package list from them. Where a lock file pins the resolved set, transitive dependencies are accurate; where only a manifest with version ranges exists, the picture is thinner. What it therefore cannot see: source vendored by copying it into the tree, anything a build script downloads, packages that only the build tool can resolve, and dependencies that live on a branch other than the default one. Container base-image contents are not there either. The supported fix for the build-time gap is the **dependency submission API** (`POST /repos/{owner}/{repo}/dependency-graph/snapshots`), which lets a workflow run the real build and submit the dependency set it resolved. This matters because the graph is upstream of everything else: vulnerability alerting, the SBOM export, and `actions/dependency-review-action` on pull requests all read it.
code
yaml · 16 linesname: Dependency review
on:
pull_request:
permissions:
contents: read
jobs:
review:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/dependency-review-action@v4
with:
fail-on-severity: highgo deeper
Recall that the graph comes from manifest and lock files committed to the repository, so a dependency nobody declared in a file is not in it.
Explain why lock-file ecosystems give accurate transitive data while manifest-only ones do not, and name the dependency submission API as the fix for build-time resolution.
Show that a graph gap silently degrades alerting, the SBOM export and pull-request dependency review at once, and say how you would detect that gap.
Own the inventory model across a portfolio: what the graph can be trusted to answer, where submission workflows are mandatory, and how you notice when one has been broken for weeks.
## How the graph is built GitHub's dependency graph is produced by parsing files, not by running anything. When a supported manifest or lock file changes on the default branch, GitHub re-parses it and updates the repository's package list. There is no build, no network capture, no bytecode analysis. Support is per-ecosystem, and the fidelity depends on what the ecosystem's files express: - **Lock files** — `package-lock.json`, `go.sum`, `Cargo.lock` and similar record the exact resolved set, including transitive packages and their pinned versions. Here the graph is close to ground truth. - **Manifests only** — a `pom.xml` or a `requirements.txt` with ranges declares intent, not resolution. GitHub gets the declared dependencies; transitive closure and the exact versions a build would choose are, by construction, less certain. This distinction is the whole answer to "why does our graph look incomplete?" in most interviews. ## What the graph is upstream of The graph is not a display feature. Three things read it: 1. **Vulnerability alerting** — matching packages and versions against advisories. 2. **The SBOM export** — the SPDX document is the graph, serialised. 3. **Dependency review on pull requests** — `actions/dependency-review-action` compares the graph for the base and head of a pull request through the dependency-graph compare endpoint, and can fail the check when the diff introduces a dependency with a vulnerability at or above a chosen severity, or a disallowed licence. So a gap in the graph is a gap in all three at once, silently. Nothing turns red to tell you a package is invisible. ## The blind spots, precisely - **Vendored code.** A library copied into `third_party/` is just source files. No manifest entry declares it, so the graph has no idea it exists — and neither does alerting. - **Build-time resolution.** Anything a build script downloads, a plugin injects, or a tool resolves without writing it to a committed file. - **Non-default branches.** The graph tracks the default branch. A long-lived release branch pinning different versions is not represented. - **Container images.** Operating-system packages in a base image are outside the repository's manifests entirely. - **Disabled graph.** On private repositories the graph is a setting. Off means empty — and empty looks exactly like clean. ## Closing the build-time gap: the dependency submission API GitHub provides a supported path for dependencies only the build knows: `POST /repos/{owner}/{repo}/dependency-graph/snapshots`. A workflow runs the real resolution — the actual build, with its actual settings — and submits the resulting dependency set as a snapshot attached to a commit and ref. Submitted dependencies join the graph and are treated like parsed ones: they feed alerting and appear in the SBOM export. This is the right answer for ecosystems where a manifest does not determine the outcome, and for builds with private mirrors, profiles or code generation. The trade-off is that the snapshot is only as current as the last workflow run, so a repository whose submission workflow has been broken for a month has a graph that quietly lags. ## Reading the graph honestly The mature framing: the dependency graph is an **inventory of declared dependencies on the default branch**, refreshed when files change, extensible by submission. It is excellent for portfolio questions — who declares this package, at what version — and it is not a description of what is running in production. Teams that conflate the two get an unpleasant surprise the first time a widely publicised vulnerability lands in something they vendored rather than declared. ## What interviewers listen for A candidate who says "GitHub scans your dependencies" has a magic-box model. A candidate who says "it parses committed manifests and lock files, transitively where a lock file exists, and there is a submission API for what the build resolves" has a mechanism model — and can therefore predict, without checking, which dependencies will be missing and why.
- Why does the graph look accurate for one repository and thin for another in the same organisation?Ecosystem, mostly. A repository with a committed lock file gives GitHub the exact resolved transitive set. A repository whose ecosystem expresses only ranges in a manifest gives it declared intent, so transitive packages and exact versions are missing until a submission workflow supplies them.
- What breaks quietly when the dependency graph is incomplete?Everything downstream, without a signal. Vulnerability matching cannot flag a package it does not know about, the SBOM export omits it, and pull-request dependency review compares an incomplete base to an incomplete head. Nothing turns red — the repository simply looks clean, which is the dangerous part.
- How current is a submitted dependency snapshot?Only as current as the workflow that submitted it. If the submission job has been failing for a month, the graph still shows last month's resolution and looks perfectly healthy. Monitor that workflow the way you would monitor any other data-freshness dependency.
saying these in an interview costs you the question
- Believing GitHub builds the project to resolve dependencies
- Assuming vendored source appears in the graph
- Thinking the graph only ever records direct dependencies
- Expecting non-default branches to be represented
- Treating an empty graph on a private repo as clean