What does it mean to 'scan a container image for vulnerabilities', and how does a tool like Trivy or Grype actually find CVEs in an image?
answer
- inventory OS packages + app lockfiles
- resolve name@version (purl)
- match NVD + distro advisories + GHSA
- reports CVE + severity + fixed version
- present != reachable; known CVEs only
basics
~20 sThe scanner inventories everything installed in the image (OS packages per layer plus app dependencies), determines each component's name and version, then matches those against vulnerability databases (like the NVD and distro advisories). Matches are reported as CVEs with severities. It is metadata matching, not runtime analysis.
solid answer
~50 sImage scanning inspects the **contents** of a built image and reports known vulnerabilities in the software it ships. A scanner like Trivy or Grype does three things: 1. **Inventory.** It unpacks the image layers and enumerates installed components: OS packages from the package manager's database (`dpkg`, `rpm`, `apk`), and language dependencies from lockfiles/manifests (`package-lock.json`, `go.sum`, `pom.xml`, `requirements.txt`). 2. **Resolve versions.** Each component becomes a `name@version` (often expressed as a package URL / purl). 3. **Match against vuln databases.** It cross-references those versions with feeds like the **NVD**, GitHub Security Advisories, and per-distro advisories (Debian, Alpine, Red Hat), producing a list of CVEs with severity (CRITICAL/HIGH/...) and, where known, the fixed version. Crucially this is **static metadata matching**: it tells you 'this image contains openssl 3.0.1, which has CVE-X', not whether your code actually calls the vulnerable path. That is why results need triage, not blind panic.
code
bash · 8 lines# Scan an image, fail nothing yet, just report
trivy image myapp:1.4.0
# Grype does the same job
grype myapp:1.4.0
# Docker's built-in
docker scout cves myapp:1.4.0go deeper
Explain the two-step idea: inventory the installed components, then match versions against CVE databases.
Add how OS vs app components are detected, the role of distro feeds, and that findings mean 'present' not 'exploitable'.
Discuss coverage gaps (hand-installed/static binaries), triage of reachability, and where to run scans.
Frame scanning as one signal in a supply-chain program; combine with SBOMs, provenance, and minimal base images to shrink the surface being matched.
## What 'scanning' means A container image is a stack of filesystem layers containing an operating-system userland plus your application and its dependencies. Any of those pieces can carry a **known** vulnerability, a publicly catalogued flaw identified by a **CVE** (Common Vulnerabilities and Exposures) id like `CVE-2024-12345`. Image scanning is the automated process of listing what is in the image and checking each piece against databases of known CVEs. It finds **known** issues only. It is not a code auditor and will not discover a novel bug in your own source; it answers 'am I shipping a version of something that the world already knows is vulnerable?' ## How a scanner works, step by step 1. **Read the image.** Tools like Trivy, Grype, and Docker Scout pull or read the image (from a registry, a tarball, or the local daemon) and walk its layers. 2. **Build an inventory.** They detect installed software two ways: - **OS packages:** by reading the package database the distro's manager maintains, `/var/lib/dpkg/status` (Debian/Ubuntu), the `rpm` DB (RHEL/Fedora), `/lib/apk/db/installed` (Alpine). This yields exact package names and versions. - **Application dependencies:** by parsing lockfiles and manifests baked into the image, `package-lock.json`/`yarn.lock`, `go.mod`/`go.sum`, `pom.xml`, `Gemfile.lock`, `requirements.txt`/`poetry.lock`, etc. 3. **Normalise identities.** Each component becomes a canonical identifier, commonly a **purl** (package URL) plus its version, so it can be matched reliably. 4. **Match against vulnerability feeds.** The scanner maintains a local copy of vulnerability data assembled from sources such as the **NVD** (US National Vulnerability Database), **GitHub Security Advisories**, and distribution advisories (Debian Security Tracker, Alpine secdb, Red Hat OVAL). Distro feeds matter because they track backported fixes: a Debian package may have a patched CVE at a version the NVD alone would still flag. 5. **Emit findings.** Output lists each CVE, the affected component and version, a **severity** (typically CRITICAL/HIGH/MEDIUM/LOW/UNKNOWN, often derived from CVSS), and the **fixed version** if one exists (or 'won't fix'/no fix). ## Why it is matching, not proof Because the technique is 'installed version X matches advisory for X', a finding means the vulnerable code is **present**, not that it is **reachable** or exploitable in your usage. A CVE in a TLS library you never invoke on an untrusted path is real but low-urgency. This gap is why scanning output is a starting point for triage, and why teams set severity policies and allow-lists rather than demanding zero findings (covered in the CI-gating question). ## What scanners do and don't cover - **Do:** OS packages, many language ecosystems, sometimes known-bad secrets or misconfigurations as add-on checks. - **Miss / weak:** software installed by hand (curl|tar into `/opt`) with no package metadata, statically-compiled binaries with vendored deps that leave no lockfile, and anything the feeds haven't catalogued yet (zero-days, or a CVE published after your DB snapshot). ## Where it runs Scanning is cheap enough to run in several places: locally before pushing, in CI on every build, and continuously in the registry so already-published images get re-flagged when **new** CVEs land against components they already contain. The next questions cover SBOMs (making the inventory a durable artifact), rebuild cadence, and failing CI on thresholds.
- Does a CVE finding mean the image is exploitable?Not necessarily. Scanners report that a vulnerable *version* is present, not that your application reaches the vulnerable code path. A CVE in a library function you never call is real but often low-urgency. Findings need triage against reachability and exposure, which is why teams gate on severity and maintain allow-lists rather than demanding zero findings.
- Why do scanners use distro-specific advisory feeds instead of only the NVD?Distributions backport security fixes without changing the upstream version number, and they mark packages 'won't fix' for some CVEs. Distro feeds (Debian, Alpine, Red Hat) capture that, so the scanner reports the CVE as fixed at the distro's patched version and avoids false positives the raw NVD would produce.
saying these in an interview costs you the question
- Thinking a scanner finds novel bugs in your own code (it finds known CVEs only)
- Believing a CVE finding automatically means exploitable
- Assuming it catches hand-installed software with no package metadata
- Confusing severity (CVSS) with exploitability in your context
- Expecting one scan to stay accurate forever as new CVEs are published