You're wiring image scanning into a CI pipeline. How do you decide which findings should fail the build, and how do you handle CVEs that have no fix yet or that you've assessed as not applicable?
answer
- two dials: severity + fixability
- fail CRITICAL/HIGH, --ignore-unfixed, --exit-code 1
- allow-list per-CVE with reason+owner+expiry (VEX)
- never blanket-ignore by severity
- scan by digest; re-scan continuously post-merge
basics
~20 sGate on severity thresholds plus fixability: typically fail on CRITICAL/HIGH that have a fix available, and don't block on unfixable ones (report them). Handle noise with a reviewed, time-boxed allow-list (e.g. .trivyignore / VEX) that documents why each CVE is ignored and expires, so exceptions get revisited rather than becoming permanent.
solid answer
~50 sTwo dials decide a fail: **severity** and **fixability**. - **Severity threshold:** commonly fail the build on `CRITICAL` and `HIGH`, warn on `MEDIUM/LOW`. Trivy: `--severity CRITICAL,HIGH --exit-code 1`. - **Fixability:** blocking on a CVE with **no available fix** just wedges delivery for something you can't act on. Add `--ignore-unfixed` so the gate fails only on findings you can actually remediate by bumping a version; still report the unfixable ones for visibility. For the inevitable false positives and 'present but not reachable' cases, use a **reviewed allow-list**: `.trivyignore`, a policy file, or **VEX** (Vulnerability Exploitability eXchange) statements that assert 'not affected' with a justification. The discipline that keeps this honest: every ignore entry has an **owner, a reason, and an expiry**, so it is revisited, not forgotten. Never blanket-ignore by severity. Operationally: scan on every build and gate there, but also **re-scan continuously** so a newly-disclosed CVE against an already-merged image raises an alert (or breaks the next build) even though nothing in the repo changed.
code
bash · 4 lines# Baseline gate: fail only on fixable CRITICAL/HIGH
trivy image --severity CRITICAL,HIGH --ignore-unfixed \
--exit-code 1 --format sarif --output trivy.sarif \
registry.example.com/myapp@sha256:abcd...go deeper
Know you gate on severity (fail CRITICAL/HIGH) via the scanner's exit code.
Add fixability (--ignore-unfixed) and the idea of a documented ignore file for false positives.
Design the full policy: severity+fixability, VEX/allow-list governance with expiry, scan-by-digest, and continuous re-scanning.
Own the org tradeoff between delivery velocity and risk: exception governance, SLAs to remediate by severity, and fleet-wide continuous scanning tied to SBOMs and provenance.
## The goal: a gate that blocks the right things A scanner will report dozens of CVEs on any non-trivial image. If the gate fails on **all** of them, developers will disable or route around it; if it fails on **none**, it is theatre. Good gating fails on findings that are both **serious** and **actionable**, and provides a disciplined escape hatch for the rest. ## Dial 1: severity threshold Severity usually derives from **CVSS** and is bucketed CRITICAL / HIGH / MEDIUM / LOW / UNKNOWN. A typical policy: - **Fail** on CRITICAL and HIGH. - **Warn/record** on MEDIUM and below. With Trivy: `trivy image --severity CRITICAL,HIGH --exit-code 1 myapp:sha`. A non-zero exit code is what actually fails the CI step. Grype uses `--fail-on high`; Docker Scout has policy equivalents. ## Dial 2: fixability A CVE with **no released fix** ('won't fix' or awaiting an upstream patch) cannot be resolved by you right now. Blocking the pipeline on it punishes teams for something outside their control and trains them to ignore the gate. So a common and sensible rule is: **fail only on fixable findings**, report the unfixable ones. - Trivy: `--ignore-unfixed`. - The unfixable set still goes into reports/dashboards and the continuous re-scan, so when a fix lands you act. Combined baseline gate: `--severity CRITICAL,HIGH --ignore-unfixed --exit-code 1`. ## The escape hatch: allow-lists and VEX, done responsibly Some fixable CRITICAL/HIGH findings still should not block, because after triage you've determined they are **not exploitable in your context** (the vulnerable function is never called, the affected feature is disabled, the component isn't in the runtime path) or the fix is a false-positive version match. Options: - **Ignore file** (`.trivyignore`, Grype config): list specific CVE ids to suppress. - **VEX (Vulnerability Exploitability eXchange):** a machine-readable statement asserting a status like `not_affected` with a justification (e.g. 'vulnerable_code_not_in_execute_path'). Scanners increasingly consume VEX so the suppression is auditable and portable rather than an opaque ignore line. The governance that prevents rot, and that interviewers look for: 1. **Per-CVE, never per-severity.** Suppress `CVE-2024-1234`, never 'all HIGH'. 2. **Documented reason + owner.** Why is it safe to ignore, and who decided? 3. **Expiry / review date.** Entries auto-expire or are re-reviewed each quarter, so 'temporary' ignores don't become permanent debt. 4. **Visible.** Ignored findings are still counted somewhere, so the risk isn't invisible. ## Fail-the-build mechanics The gate is just an exit code plus policy. Keep it fast and deterministic: - Pin/cache the vuln DB for reproducibility within a run, but refresh regularly. - Scan the **exact** image you'll ship (by digest), after build, before push/deploy. - Emit machine-readable output (SARIF/JSON) so results appear in the PR and security dashboards, not just logs. ## Don't stop at merge: continuous re-scanning Critically, a build-time gate only sees today's CVEs. Tomorrow a new CVE may be disclosed against a component already merged and deployed. So pair the gate with a **scheduled re-scan** of stored images/SBOMs (see the SBOM and cadence questions) that can open a ticket, alert, or fail the next scheduled build. Otherwise 'green at merge' silently rots. ## Anti-patterns to avoid - Failing on every severity, then having teams add `|| true` to the scan step, gutting the control. - A giant unreviewed ignore file with no reasons or expiry. - Blocking on unfixable CVEs, creating unactionable red builds. - Scanning a different tag than you deploy (scan by digest). - Treating the merge-time green as permanent. The crisp summary: **fail on serious AND fixable; suppress only per-CVE with reason, owner, and expiry; and re-scan continuously so new disclosures aren't missed.**
- Why exclude unfixed CVEs from the failing gate instead of blocking on them?Because there is no version to bump to, blocking the build makes it red for something the team cannot remediate, which trains people to bypass the gate. You still report and track unfixable CVEs and re-scan continuously, so the moment a fix is released the gate can start enforcing it and you rebuild.
- How do you stop an allow-list from becoming a permanent dumping ground?Require every entry to be per-specific-CVE, carry a documented justification and an owner, and have an expiry or scheduled review. Entries that expire force a re-triage. Blanket suppressions by severity are banned, and ignored findings still appear in reports so the accepted risk stays visible.
- A pipeline was green at merge but the image is now flagged critical in production. How is that possible and what prevents it?New CVEs are disclosed after merge against components already shipped, so a one-time build gate can't catch them. The fix is continuous re-scanning of deployed images or their stored SBOMs against fresh feeds, which alerts or fails the next build so you rebuild on the patched base.
saying these in an interview costs you the question
- Failing on all severities, prompting developers to disable the scan step
- Blocking builds on CVEs that have no available fix
- Maintaining an unreviewed, undated blanket ignore list
- Suppressing by severity ('ignore all MEDIUM') instead of per-CVE with justification
- Believing a green scan at merge stays valid forever (no continuous re-scan)
- Scanning a different tag than the digest actually deployed