When a Trivy 0.74 scan logs 'Downloading vulnerability DB', what is it fetching, from where, and when will it fetch again?
answer
- an OCI artifact, not an image
- two registries tried in order
- trivy.db plus metadata.json
- NextUpdate and the one-hour rule
basics
~10 sIt pulls trivy-db, a BoltDB file of compiled advisories shipped as an OCI artifact, from mirror.gcr.io/aquasec/trivy-db:2 and then ghcr.io/aquasecurity/trivy-db:2, caches it, and downloads again once its metadata says a newer build is due.
solid answer
~40 sTrivy ships the scanning engine but not the advisory data. On a scan that needs vulnerability matching it pulls `trivy-db` — distro and language advisories compiled into a BoltDB file — as an OCI artifact. In 0.74 the default `--db-repository` list is `mirror.gcr.io/aquasec/trivy-db:2`, then `ghcr.io/aquasecurity/trivy-db:2`; the `:2` tag is the DB schema version, not a date. The files land in the cache directory under `db/` as `trivy.db` and `metadata.json`. On later runs Trivy reads `metadata.json` and skips the download while the clock is before `NextUpdate`, or if the DB was downloaded within the last hour. A separate `trivy-java-db` — an index of JAR artifacts, not advisories — is pulled only when a scan finds a JAR.
code
bash · 3 linestrivy image --download-db-only
trivy version --format json
trivy image --db-repository mirror.gcr.io/aquasec/trivy-db:2 --db-repository ghcr.io/aquasecurity/trivy-db:2 registry.example.internal/payments/api:1.8.2go deeper
Recall that trivy-db is a separate download, pulled as an OCI artifact into the cache directory, and that the scan results depend on how fresh it is.
Explain the ordered repository list, the schema-version tag, the trivy.db and metadata.json pair, and the NextUpdate and one-hour rules that decide whether a scan downloads again.
Show you can tell a stale-DB discrepancy from a real change, read UpdatedAt from trivy version, and explain why a denied registry error does not fall back while a 429 does.
Frame the database as a dependency of every scan result: who owns its freshness, its registry availability and its schema compatibility across Trivy upgrades in a fleet of runners.
## What the download actually is A Trivy binary contains the **scanning engine**: the analysers that find OS packages, lockfiles and binaries in a target, and the logic that compares versions. It does not contain the **advisory data** those versions are compared against. That data is `trivy-db`, a single **BoltDB** file (`trivy.db`) into which advisories from distribution security trackers and language advisory databases have already been compiled, plus a small `metadata.json` that describes the build. The database is published as an **OCI artifact** — it travels through ordinary container registries, but it is not a runnable image. Its layer uses its own media type, `application/vnd.aquasec.trivy.db.layer.v1.tar+gzip`, which matters later if you proxy or mirror it. ## Where Trivy 0.74 looks for it Trivy keeps a list of repositories and tries them **in order of priority**. In 0.74 the defaults are: | Database | First repository | Second repository | |---|---|---| | `trivy-db` (advisories) | `mirror.gcr.io/aquasec/trivy-db:2` | `ghcr.io/aquasecurity/trivy-db:2` | | `trivy-java-db` (JAR index) | `mirror.gcr.io/aquasec/trivy-java-db:1` | `ghcr.io/aquasecurity/trivy-java-db:1` | Points that confident answers get wrong: - The tag is the **schema version** of the database format, not a date or a build number. If you pass a repository without a tag, Trivy appends the schema number itself and logs that it is doing so for backward compatibility; it never asks for `latest`. - Moving to the second repository is not automatic on any failure. Trivy tries the next one only on a **temporary** registry error (rate limiting with status 429, a 5xx) or a `BLOB_UNKNOWN` error. A denied or unauthorised response stops the download with an error and a pointer to the troubleshooting page. - Several repositories per flag have been supported since 0.56.0, and the same list can be set with `--db-repository` and `--java-db-repository` or their `TRIVY_`-prefixed environment variables. ## Where it lands The artifact is unpacked into the **cache directory** (`--cache-dir`), in a `db/` subdirectory holding `trivy.db` and `metadata.json`. The Java index lands in `java-db/` as `trivy-java.db` with its own `metadata.json`. `trivy clean --vuln-db` or `trivy clean --java-db` removes them; `trivy version --format json` prints the cached DB's `UpdatedAt` and `NextUpdate`. ## When it downloads again On every scan Trivy decides whether the cached copy is good enough, in roughly this order: 1. If `trivy.db` or `metadata.json` is missing, it must download — and `--skip-db-update` on that first run is an error, not a shortcut. 2. If the cached schema is newer than the binary understands, it stops and tells you to upgrade Trivy; if the schema differs otherwise, it downloads a matching one. 3. If the current time is still before the `NextUpdate` recorded in `metadata.json`, it keeps the cached copy. 4. If the cached copy was downloaded within the **last hour** (`DownloadedAt`), it keeps it even when `NextUpdate` has passed. 5. Otherwise it downloads a fresh build. So a busy runner does not hit the registry on every scan, and a runner that has been idle for a day refreshes on its next scan. ## The Java DB is a different animal `trivy-java-db` contains **no CVEs**. It is an index that lets Trivy identify a JAR — its group, artifact and version — from the JAR's SHA-1 digest when the JAR's own `pom.properties` and `MANIFEST.MF` do not say enough. Java advisories themselves sit in `trivy-db`, sourced from the GitHub Advisory Database. Trivy downloads the Java index lazily, only when a scan actually finds a JAR, which is why a first Java-heavy scan shows a second download step. A third artifact, the misconfiguration checks bundle, follows the same OCI pattern but belongs to IaC scanning, and it is also embedded in the binary as a fallback; the vulnerability database is not. ## Why an operator cares - A scan result is only as current as `UpdatedAt`. Two runs of the same image a week apart can disagree purely because the database moved. - `trivy image --download-db-only` warms the cache without scanning, which is how CI jobs and offline transfers prepare a known copy. - Registry rate limits are the classic cause of a failed first step on shared runners; the ordered repository list exists to ride them out. ## Diagnosing a failed download step When the first step fails, the error text usually names the cause: - **`failed to download vulnerability DB`** behind a corporate firewall: the runner cannot reach the registry hosts. For the defaults these are `mirror.gcr.io` and `googlecode.l.googleusercontent.com`, then `ghcr.io` and `pkg-containers.githubusercontent.com`; all must be allowed, or the DB must come from an internal mirror. - **`DENIED: denied` from `ghcr.io`**: a stale GitHub token in the environment or in Docker's stored credentials is being sent with the pull. Trivy's troubleshooting guide says to remove it (`docker logout ghcr.io` or unset `GITHUB_TOKEN`) and retry; this error does not fail over to another repository. - **A 429 on the first repository followed by success on the second** is the fallback working as designed, not a fault. - **Schema errors after an upgrade** mean the cached file and the binary disagree about format; let Trivy download a matching build rather than copying an old file back in.
- Why does a --db-repository value without a tag still pull the :2 tag?Trivy parses the reference and, when no tag is given, appends the DB schema version rather than `latest`, logging that it is adding the schema version for backward compatibility. The tag names the database format a given Trivy release can read, so a mirror must keep the schema tag the binary expects.
- What happens after a Trivy upgrade if the cached DB uses a different schema?A schema mismatch counts as needing an update, so a normal scan downloads a matching build. If the cached schema is newer than the binary supports, Trivy stops and says the Trivy version is old. With `--skip-db-update` and an old schema, the scan fails with an error saying the flag cannot be used with the old DB schema.
saying these in an interview costs you the question
- The vulnerability database is embedded in the Trivy binary, so no download is needed
- The :2 tag on trivy-db is the date or build number of the snapshot
- Trivy downloads the full vulnerability database on every scan
- If the first registry denies access, Trivy quietly falls back to the next one
- trivy-java-db holds the CVE records for Maven packages